About
A research project turned into a working screening tool
This platform records a short video, scores it with three trained neural networks, and turns the result into four affective indicators. Every figure published on the site is transcribed from the training notebook — nothing is estimated.
- Modalities
- 2 Modalities
- Emotion classes
- 8 Emotion classes
- Training samples
- 2,452 Training samples
What it is
One recording, two modalities, one fused verdict
A user answers five guided prompts on camera. The audio track and the video frames are scored separately — a CNN/BiLSTM network over 128-band log-Mel spectrograms for speech, a ResNet3D network over 16-frame clips for facial motion — and their features are concatenated into a single 568-dimension vector that a fusion network turns into an emotion distribution.
A transparent rule layer maps that distribution onto four indicators: possible depression, anxiety, stress and emotional blunting. There is no second model and no learned weights in that layer, so every score can be traced back to the emotion probabilities that produced it.
Why it exists
Screening is slow, self-reports are noisy
Standard questionnaires rely on a person recognising and reporting their own state. That is exactly what a depressive episode interferes with, and it makes early screening late, inconsistent and easy to skip.
Voice and facial motion carry affective signal that does not depend on self-report. Reading both together gives a second, objective view that a clinician, counsellor or the user can weigh alongside everything else they already know.
How it is built
A web application with a machine-learning sidecar
The web layer never runs a model. A submitted check is queued, and a Python sidecar loads the trained networks, runs the pipeline and hands the result back.
Web application
Server-rendered pages, an SPA-style navigation layer and a bilingual RTL/LTR interface for both the public site and the admin dashboard.
Laravel 13 · Livewire 4 · Tailwind CSS 4
Processing queue
Every check is queued rather than processed in the request, so a slow inference run never holds a browser connection open.
Laravel queues · database driver
Inference sidecar
A separate Python process loads the audio, video and fusion networks exactly as the training notebook produced them, and returns the full probability distribution.
Python sidecar · PyTorch · TensorFlow
Storage
Recordings sit on a private disk tied to the account that produced them. Results, scores and the model breakdown are stored alongside the check.
MySQL · media library · private disk
Audio
CNN + BiLSTM + RNN over log-Mel spectrograms
Video
ResNet3D r3d_18 over 16-frame clips
Fusion
MLP over the 568-dim fused vector
Principles
What we hold ourselves to
An emotional-state model is only usable if the people it reads know what happens to their recording, and what the number means.
Consent first
Nothing is analysed without a deliberate submission. There is no passive listening, no background capture and no analysis of a recording a user did not send.
Minimal data
We ask for what a check needs and nothing more. A recording belongs to the account that produced it, and deleting the check deletes the file with it.
Explainable output
Every result ships with the emotion distribution, the per-stage predictions and the rule that produced each indicator — never a bare verdict.
Not a diagnosis
The output is a screening signal. It is worded as a possibility, it is never presented as a clinical conclusion, and it always points back to a professional.
Limitations
What this model cannot do
The model was evaluated on a held-out split of the training dataset. Real recordings differ from it in ways that matter, and the honest reading of a result depends on knowing how.
-
Trained on acted emotion
The dataset is acted emotional speech and song. Acted affect is more pronounced than everyday affect, so real-world performance is expected to be lower than the held-out figures.
-
English speech only
The speech network was trained on English. Answers spoken in another language degrade the audio stage, and with it the fused result.
-
No context
The model sees two minutes of a person. It knows nothing about their history, medication, circumstances or what happened that morning — all of which a clinician would weigh.
-
Screening, not assessment
Indicators are prompts to look closer, not conclusions. They are not validated as a diagnostic instrument and must not be used to make a clinical or employment decision on their own.
This result is not a medical diagnosis. Please consult a professional.
See how it fits your case
See what a check gives you, or write to the team with a question.